Papers with cross-lingual memorization

2 papers
Shared Path: Unraveling Memorization in Multilingual LLMs through Language Similarities (2025.emnlp-main)

Copied to clipboard

Challenge: Using multilingual models, we find that treating languages in isolation obscures the true patterns of memorization.
Approach: They propose a graph-based correlation metric that incorporates language similarity to analyze cross-lingual memorization.
Outcome: The proposed model incorporates language similarity to analyze cross-lingual memorization in 95 languages.
OWL: Probing Cross-Lingual Recall of Memorized Texts via World Literature (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are known to memorize and recall English text from their pretraining data, but the extent to which this ability generalizes to non-English languages or transfers across languages remains unclear.
Approach: They propose a dataset of 31.5K aligned excerpts from 20 books in ten languages, including English originals, official translations and new translations in six low-resource languages.
Outcome: The proposed model can recall English content in translations, but perturbations reduce performance, causing the model to fail.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations